I run an M1 MacBook Pro with 32GB of RAM, and over time I've put together a workflow that combines macOS, a Kali Linux VM, and several AI coding assistants — some running locally, some rented on demand. This article covers the full setup: how the pieces fit together, and the exact config steps to reproduce each one.
My main security-lab environment — Burp Suite, PortSwigger labs, malware analysis — lives in a Kali Linux VM. I went through a few dead ends before landing on the current setup.
I first tried Docker for a Kali GUI, but containers have no display server by default on macOS, so getting a desktop meant fighting with XQuartz or a VNC bridge just to see a screen. I then tried UTM with QEMU, which is the more Apple-Silicon-native path, but hit the usual wall: Kali's official virtual machine images are built for AMD64, so ARM64 support on M1 hardware is inconsistent and occasionally unstable.
VMware Fusion ended up being the most stable option for daily use. VMs built this way are also portable: moving to a new Mac later just means copying the VM package and reopening it, no reinstall required.
| Approach | Result on M1 |
|---|---|
| Docker (Kali image) | No display server by default — needs XQuartz/VNC workaround |
| UTM + QEMU (Kali ARM64) | Official images mostly AMD64 — inconsistent support |
| VMware Fusion | Stable, portable, my daily driver |
Instead of installing AI models twice, I run them on the Mac itself with Ollama and reach them from inside the Kali VM over the LAN.
Ollama binds to localhost by default. To make it reachable from the VM, I start it temporarily in the foreground:
OLLAMA_HOST=0.0.0.0:11434 ollama serve
This is intentionally a foreground-only command — hitting Ctrl+C reverts the binding completely, with no persistent launchctl environment variable left exposed on the network.
Find the Mac's LAN IP to use in both CLI configs below:
ipconfig getifaddr en0
| Model | Size | Best for |
|---|---|---|
qwen3.5:9b-mlx | ~8.9GB | Daily driver — fast, RAM-efficient coding & reasoning |
gemma4:12b-mxfp8 | ~13GB | Multimodal / vision tasks |
gemma4:26b-mlx | ~17GB | Stronger multimodal reasoning |
igorls/gemma-4-12B-it-heretic-GGUF:Q8_0 | ~12GB | Community uncensored fine-tune |
Create the config directory and file inside the VM:
mkdir -p ~/.qwen
nano ~/.qwen/settings.json
Point it at Ollama's OpenAI-compatible endpoint on the Mac:
{
"env": {
"OLLAMA_API_KEY": "ollama"
},
"modelProviders": {
"openai": [
{
"id": "qwen3.5:9b-mlx",
"name": "Qwen3.5 9B (Mac)",
"baseUrl": "http://<MAC_LAN_IP>:11434/v1",
"envKey": "OLLAMA_API_KEY"
}
]
},
"model": {
"name": "qwen3.5:9b-mlx"
}
}
Ollama doesn't validate API keys locally, so the placeholder value "ollama" is enough to satisfy Qwen CLI's envKey requirement.
Validate the JSON before relying on it:
python3 -m json.tool ~/.qwen/settings.json
Run it:
qwen
Switch models mid-session with /model.
cn)Install it if needed:
sudo apt install -y nodejs npm
npm install -g @continuedev/cli
cn --version
Create the config file:
mkdir -p ~/.continue
nano ~/.continue/config.yaml
name: Local Config
version: 0.0.1
schema: v1
models:
- name: Qwen3.5 9B (Mac)
provider: ollama
model: qwen3.5:9b-mlx
apiBase: http://<MAC_LAN_IP>:11434
roles:
- chat
- edit
- apply
Additional models can be added the same way — one entry per model — so all of them are selectable inside a session.
Continue CLI auto-detects ~/.continue/config.yaml, so no flag is needed:
cn
If it doesn't pick the file up automatically:
cn --config ~/.continue/config.yaml
Note: Continue may show a tool-calling capability warning for some models (e.g. gemma4:12b-mxfp8) — that's Continue's internal compatibility check, not a hard failure.
For models larger than 32GB of unified memory can comfortably run — like a 31B heretic fine-tune — I rent GPU compute on Vast.ai instead of running it locally.
On the Vast.ai instance:
ollama pull tinyrick/gemma-4-31B-it-uncensored-heretic-vision-llmfan46:Q4_K_M
ollama list
Expose it through the instance's tunnel (typically Cloudflare), then point either CLI config at the tunnel URL instead of the Mac's LAN address:
{
"id": "tinyrick/gemma-4-31B-it-uncensored-heretic-vision-llmfan46:Q4_K_M",
"name": "Gemma 4 31B Heretic (Vast.ai)",
"baseUrl": "https://<your-tunnel-url>/v1",
"envKey": "OLLAMA_API_KEY"
}
Sanity-check the endpoint before trusting it:
curl https://<your-tunnel-url>/v1/chat/completions \
-H "Authorization: Bearer ollama" \
-H "Content-Type: application/json" \
-d '{"model": "tinyrick/gemma-4-31B-it-uncensored-heretic-vision-llmfan46:Q4_K_M", "messages": [{"role": "user", "content": "Say hi in one word"}]}'
A clean response means the model is reachable and ready to select inside qwen or cn. Vast.ai instances run in unprivileged Docker containers isolated per tenant, which gives a reasonable degree of technical separation — worth factoring in if prompt privacy matters, since it remains a shared physical machine.
| Trade-off | Detail |
|---|---|
| Full control | No subscriptions, no vendor lock-in, every component swappable |
| Free (mostly) | Only cost is on-demand GPU rental for the bigger models |
| More moving parts | Three layers (VM, host, remote) means more to keep track of than a single app |
| Easy to lose track | IP addresses and tunnel URLs change — worth keeping a running note of the current ones |
This keeps the security tooling isolated inside the VM while still giving full AI assistance, switching between a fast local model and a larger remote one depending on what the task needs.